Formatting: introduce esc_attr_name() for HTML attribute name sanitization - #12915
Formatting: introduce esc_attr_name() for HTML attribute name sanitization#12915SainathPoojary wants to merge 2 commits into
Conversation
Test using WordPress PlaygroundThe changes in this pull request can previewed and tested using a WordPress Playground instance. WordPress Playground is an experimental project that creates a full WordPress instance entirely within the browser. Some things to be aware of
For more details about these limitations and more, check out the Limitations page in the WordPress Playground documentation. |
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the Core Committers: Use this line as a base for the props when committing in SVN: To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
irozum
left a comment
There was a problem hiding this comment.
Nice addition — esc_attr_name() fills a real gap since esc_attr() only ever escaped values, not names, and the allowlist regex correctly strips all the characters that would let an attacker break out of the attribute-name context (quotes, =, <, >, /, whitespace, control chars). I checked out the branch, ran the new test file (33/33 pass) plus composer lint:errors and composer phpstan — both clean on the changed files (the pre-existing lint failures are in unrelated files).
One functional issue: src/wp-includes/formatting.php:20 (preg_replace( '/[^a-zA-Z0-9_.:\[\]-]+/u', '', $safe_text )) can return null instead of a string. wp_check_invalid_utf8( $text ) is called with the default $strip = false, so on a site whose blog_charset isn't UTF-8, is_utf8_charset() short-circuits it to return $text unmodified even if it contains invalid UTF-8 bytes. Those bytes then hit preg_replace() with the /u modifier, which returns null on invalid UTF-8 rather than throwing — I verified this directly (preg_replace('/[^a-zA-Z0-9_.:\[\]-]+/u', '', "caf\xE9") → NULL). That breaks the documented @return string contract. esc_attr()/esc_html() don't have this problem because _wp_specialchars() matches with a non-/u (byte-safe) pattern instead. Worth either passing wp_check_invalid_utf8( $text, true ) to always strip invalid bytes up front, or coalescing a null result from preg_replace() to ''. The test suite doesn't catch this because PHPUnit runs with the UTF-8 default charset.
Also, @since 7.1.0 on both the function and the esc_attr_name filter doc is likely wrong — trunk is currently 7.1-RC2 (feature-frozen), and other new hooks/functions already in trunk are tagged @since 7.2.0 (e.g. comment.php's recent additions), so this should probably target 7.2.0 too.
The HTML5 spec allows arbitrarily named attributes (e.g., data-*), but WordPress lacked a dedicated function to escape dynamically generated attribute names. Developers sometimes incorrectly used esc_attr() for this, which only escapes attribute values.
This PR introduces esc_attr_name() in formatting.php to resolve this. It uses a strict allowlist regex ([^a-zA-Z0-9_.:\[\]-]+/u) to strip any characters not explicitly permitted in an HTML attribute name per the HTML5 specification (such as spaces, quotes, control characters, <>, =, etc.).
Trac ticket: #43010
This Pull Request is for code review only. Please keep all other discussion in the Trac ticket. Do not merge this Pull Request. See GitHub Pull Requests for Code Review in the Core Handbook for more details.